Back

BMJ Health & Care Informatics

BMJ

Preprints posted in the last 30 days, ranked by how well they match BMJ Health & Care Informatics's content profile, based on 15 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation

Chowdhury, A. R.; Chowdhury, B.

2026-09-02 health informatics 10.64898/2026.09.01.26361908 medRxiv
Top 0.1%
12.8%
Show abstract

Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.

2
Community Learning Ledgers for Cancer Navigation in Small Island Developing States

Roach, A.; Amow, A.; Haraksingh, R.; Archer, N.; Cyrus, E.; Evans, A. N.; Calleja, N.; Croes, R.; Forghani, I.; Bajnath, A.; Hadley, D.

2026-08-18 health informatics 10.64898/2026.08.16.26360547 medRxiv
Top 0.1%
11.7%
Show abstract

Importance. Cancer is the second leading cause of death among patients in the Caribbean, where outcomes are associated with delayed clinical navigation to screening, diagnosis, and treatment. Artificial intelligence is increasingly used to guide patients with cancer to care, but whether these systems provide clinically actionable, facility-verified guidance for individuals in this population, and whether governance of the system is associated with the quality of that guidance, has not been evaluated. Objective. We tested whether a governed community learning platform navigates Caribbean cancer patients better than four ungoverned AI systems, and we tracked how community intelligence accumulates over time. Design, Setting, and Participants. We deployed a community learning ledger (CaribChat.ai) across ten Caribbean jurisdictions beginning March 2, 2026, and report all sessions through June 1, 2026 (N=207). An initial actively-promoted accrual period (March 2 - April 6, 2026; 168 sessions) was followed by continued organic use after active clinical promotion ceased. We then submitted the same 28 patient screening queries to ChatGPT (GPT-4o), Claude Haiku 4.5, DeepSeek-Chat, and OpenEvidence on April 5-6, 2026. Claude Haiku 4.5 powers CaribChat; testing it without governance isolates the governance effect. The platform requires no registration. Exempt under 45 CFR 46.104(d)(4)(ii). Main Outcomes and Measures. We classified 207 community sessions by thematic domain and temporal phase. We scored each of five systems on Caribbean facility citation, actionable navigation, and US-resource leakage across 28 screening queries. Results. The ledger accumulated 207 sessions - 168 during an actively-promoted accrual period (March 2 - April 6) and 39 after active clinical promotion ceased. Community engagement evolved from screening questions to active treatment navigation and diaspora engagement. CaribChat cited verified Caribbean facilities in 28/28 (100%) responses versus 10/28 (35.7%) for ChatGPT and 9/28 (32.1%) for OpenEvidence. CaribChat provided actionable navigation in 28/28 (100%) versus 2/28 (7.1%) for OpenEvidence (P<=.001). The same model scored 100% with governance and 54% without (P<=.001). DeepSeek cited US resources in 57.1% of Caribbean responses. After active clinical promotion ceased, off-codebook queries rose from 2.4% to 26.7% across phases while the governance contract continued to reject every adversarial probe - the community persisted but drifted from the cancer codebook absent clinician curation. The deployment operated within the OECS Health Strategy 2030 and CARICOM regional health frameworks, with queries originating across Caribbean jurisdictions led by Trinidad and Tobago. Conclusions and Relevance. Every ungoverned AI system we tested failed Caribbean cancer navigation. The best scored 68%. The most widely adopted physician platform scored 7%. The same foundation model scored 100% with governance and 54% without. Community intelligence accumulated from the population it serves, not published literature, is what makes health AI work in SIDS. The post-promotion decay shows the requirement is bidirectional: sustained, on-codebook engagement depends on patients and clinicians working together - community participation and active clinical curation are jointly necessary for maximum AI leverage.

3
Development of an interdisciplinary network to improve the capacity to conduct digital legacy research: a quality improvement initiative

Nwosu, A. C.; Tibbles, A.; Goodwin, C.; Kaye, L.; Stanley, S.

2026-08-10 palliative medicine 10.64898/2026.08.06.26359876 medRxiv
Top 0.1%
10.3%
Show abstract

Background Digital legacy (the digital information available about someone following their death) has increasing societal importance as personal assets and interactions become increasingly digitized. Healthcare professionals often have a limited understanding of how to address digital legacy in practice, and there is a lack of interdisciplinary networks to improve education, research, and professional development in digital legacy. Objective This paper describes the development of an interdisciplinary initiative designed to build research capacity and develop consensus-based recommendations for integrating digital legacy into palliative care. Method Over 12-months, we conducted interdisciplinary engagement activities with diverse stakeholders, including clinicians, designers, and sociologists. We used a modified World Cafe method to facilitate dialogue and capture feedback on how memories are digitally curated, the management of digital estates, and intergenerational perspectives on digital legacy. Results We identified eight core recommendations for research and policy, including promoting digital legacy education, supporting policy development, and broadening the scope of interdisciplinary research. Our discussions highlighted the complexity of modern digital estates and the need for legal and ethical frameworks to protect individual rights. Conclusions The Network demonstrates that interdisciplinary collaboratives can address important issues relating to digital legacy, which provides a foundation to conduct collaborative research that improves the management of digital legacies in society.

4
Reducing Under-Triage Risk in Large Language Model Based Clinical Triage Using UMLS-CUI Augmentation

Gokhale, R.; Kukreja, M.; Kumar, N.; Gourab, K.

2026-08-10 health informatics 10.64898/2026.08.07.26358932 medRxiv
Top 0.1%
6.9%
Show abstract

Background: Public facing large language models (LLMs) are increasingly used for health guidance, including triage recommendations. We evaluated whether augmenting LLM prompts with standardized clinical concepts from the Unified Medical Language System (UMLS) could improve the safety and robustness of clinical triage recommendations. Methods: We used a publicly available dataset comprising 60 clinician-authored clinical vignettes, each represented in 16 demographic and narrative variations, yielding 960 vignette-factor combinations. Clinical entities were extracted using a two-stage pipeline combining ClinicalBERT-based named entity recognition with rule-based identification of laboratory abnormalities. Extracted entities were mapped to UMLS Concept Unique Identifiers (CUIs). Negated concepts were excluded. A confidence-weighted CUI voting classifier was trained using empirical associations between CUIs and clinician-assigned triage categories. We compared five approaches: CUI-only classification, MedGemma 27B, MedGemma 27B augmented with CUIs, GPT-4o-mini, and GPT-4o-mini augmented with CUIs. Outcomes included overall accuracy, under-triage, over-triage, emergency-case accuracy, and sensitivity to anchoring statements. Results: CUI augmentation decreased under-triage but increased over-triage in both models tested (GPT-4o-mini and MedGemma 27B). It improved high-acuity recognition while reducing recognition of low-acuity cases. CUI augmentation had mixed effects on overall triage accuracy; accuracy increased for MedGemma 27B but decreased for GPT-4o-mini. Emergency-case accuracy improved from 73.0% to 80.7% for GPT-4o-mini and from 60.5% to 68.5% for MedGemma 27B. CUI augmentation also reduced susceptibility to anchoring statements. These findings suggest that the principal value of CUI augmentation may be shifting model behavior toward safety-oriented behavior rather than uniformly improving overall accuracy. Conclusion: Ontology-grounded prompt augmentation shifted LLM triage recommendations toward greater sensitivity to high-acuity presentations and reduced overall under-triage. These safety gains were accompanied by increased over-triage and mixed effects on overall accuracy. A hybrid architecture combining LLM-based language understanding with interpretable UMLS-derived clinical concepts may improve the safety and robustness of AI-assisted triage. Further evaluation using real-world patient communications and clinical outcomes is warranted.

5
Evaluating Eight Retrieval-Augmented Generation (RAG) Large Language Models' Responses to Clinical Questions: A Comparative Study

Krump, P. A.; Blasingame, M. N.; Koonce, T. Y.; Williams, A. M.; Su, J.; Giuse, N. B.

2026-08-12 health informatics 10.64898/2026.08.10.26360108 medRxiv
Top 0.1%
6.7%
Show abstract

Background: Large language models (LLMs) that use retrieval-augmented generation (RAG) are increasingly used to answer clinical questions, although the evaluation of these systems remains limited. Building on previous studies conducted by our team, this case report aimed to improve upon this knowledge gap by applying a reusable methodology to compare the performance of eight LLMs that utilize RAG techniques for evidence synthesis. Case Presentation: Eight commercially available RAG LLM tools (OpenEvidence, Undermind, Consensus, SciSpace, Elicit, MediSearch, EvidenceHunt, and Scite) were evaluated using twelve ChatGPT-generated clinical questions on the topics of treatment, etiology, and prognosis. To enable comparison, we prompted ChatGPT to identify all key unique medical concepts from the full set of LLM responses to each question. Concepts were categorized as critical ("must-have") or non-critical ("nice-to-have") for answering the clinical question. Experienced information scientists were consulted at each step for their expertise. Descriptive statistics and Kruskal-Wallis tests were used to compare performance across tools and question categories. No significant differences were found among the eight RAG LLMs in their coverage of "must-have" (p=0.95) or "nice-to-have" (p=0.16) key unique medical concepts, and no single tool consistently captured all identified concepts. Conclusions: These findings suggest that RAG LLMs may be supplementary tools for evidence retrieval and synthesis but cannot, at this time, fully replace comprehensive expert review of the medical literature. The evaluation framework presented here may be a useful model for future comparative assessments of rapidly evolving AI evidence synthesis tools.

6
Beyond Length of Stay: Patient and Carer Perspectives on Virtual Hospital Pathways Following Colorectal Surgery

Reza, L.; Arbai, Z.; Ward, H.; Payne, L.; Kinross, J.; Patel, V.

2026-08-27 surgery 10.64898/2026.08.24.26361282 medRxiv
Top 0.1%
6.6%
Show abstract

Background Virtual hospital (VH) pathways support early discharge through remote monitoring, but limited evidence has hindered implementation in colorectal surgery. This study aimed to define patient- and carer-relevant outcomes and experiences of VH following colorectal surgery. Methodology A patient and public involvement and engagement (PPIE) consultation was conducted with 8 participants (7 patients, 1 carer; 4 women, 4 men) who had experienced VH following bowel resection at a high-volume robotic unit. Purposive sampling ensured that 50% of participants had experienced readmission. The 90-minute session was delivered via Microsoft Teams. Data were analysed using reflexive thematic analysis. Results Seven themes were identified: readmission, remote monitoring, carer burden, recovery, equity, readiness for discharge, and information delivery. Patients supported early discharge when remote monitoring enabled timely detection of complications and readmission pathways were efficient. Readmission was not perceived as failure but as appropriate escalation. Dissatisfaction with readmission was related to delays in emergency care. Remote monitoring provided psychological safety, with patients feeling held at home. Carers assumed substantial, often unrecognised, quasi-clinical roles. Recovery was defined by return to function rather than length of stay. Equity concerns were evident, with VH favouring those with adequate support at home, digital literacy, and language proficiency. Discharge readiness was both clinical and psychological. Information delivery at discharge was often poorly retained and requires reinforcement preoperatively at every encounter with patients and carers. Conclusions VH pathways are acceptable and valued. Readmission is a marker of system responsiveness rather than failure of early discharge on VH. Psychological preparedness, carer support, and equitable access are critical to successful and scalable implementation of early discharge using a virtual hospital.

7
Pragmatic trial design of a digital supportive care platform for patients with brain tumours and their carers

Kalla, M.; Bray, S. C.; Schadewaldt, V.; Krishnasamy, M.; Whittle, J. R.; Chapman, W.; Huckvale, K.; Burns, K.; Capurro, D.; Layton, M. J.; Thomas, J.; Lourenco, R. D. A.; Andrew, D.; McAlpine, H.; Dhillon, R. S.; Cain, S.; Rosenthal, M.; Drummond, K. J.

2026-08-21 health informatics 10.64898/2026.08.18.26360754 medRxiv
Top 0.1%
6.6%
Show abstract

Patients with a brain tumour receive evidence-based clinical care in Australia but a focus on supportive care, including social connection, is often deficient. Digital health platforms hold promise to support these patients and their carers. Existing platforms often lack end-user co-design, evidence-based development and rigorous evaluation. Recognising this unmet need, we co-designed Brain Tumours Online, a digital supportive care platform to streamline access to educational resources, symptom management tools, and peer support for patients, carers, and healthcare professionals. In this article, we present our evaluation approach for Brain Tumours Online to advance methodological thinking in the evaluation of multi-faceted, co-designed digital health platforms. In contrast to standardised procedures in clinical trials, digital health interventions such as supportive care platforms are more complex due to their interactive nature, no prescriptive protocols for usage and the dynamic content of web-based information. Thus, traditional evaluation approaches often fall short in evaluating such multi-faceted digital health supportive care platforms. To address these challenges, we developed a bespoke, logic-modelling based evaluation approach to assess the usability, engagement, impact, and economic value of our platform. Our pragmatic but rigourous evaluation approach required the adaptation of existing evaluation frameworks, subject-matter, and lived experience expert knowledge. Our implementation science and co-design approach are shared in different papers. Our study outcomes will also be shared in a separate paper. In the current paper, we share our approach to the evaluation of Brain Tumours Online and provide insights that may be of value for other researchers interested in the nuances of trialing multi-faceted digital health supportive care platforms.

8
Accuracy and error patterns of ChatGPT-4o for real-time English-Nepali voice translation: A cross-sectional field evaluation in rural Nepal

Mandich, A.; Koirala, S.; Westen, S.; Adhikari, S.; Acharya, A.; Shrestha, A.

2026-08-28 health informatics 10.64898/2026.08.25.26361303 medRxiv
Top 0.2%
5.7%
Show abstract

Language discordance can impede community-based research and health communication where trained interpreters are limited. Although multimodal artificial intelligence systems can provide real-time spoken translation, performance with under-resourced languages during spontaneous field interactions remains poorly characterized. We evaluated ChatGPT-4o during bidirectional English-Nepali voice translation in a community setting near Dhulikhel Hospital, Nepal. In this cross-sectional field study, 30 primarily Nepali-speaking adults were recruited by convenience sampling. ChatGPT-4o mediated conversations using standardized English questions and spontaneous Nepali responses. A bilingual Nepali-English reviewer assessed 485 translated utterances using a 3-point accuracy scale and an inductively developed framework for translation and conversational deviations. Of 485 translations, 282 (58.1%) received the highest accuracy rating, 134 (27.6%) a moderate rating, and 69 (14.2%) the lowest. Mean accuracy was higher for English-to-Nepali than Nepali-to-English translation (2.63 {+/-} 0.53 vs 2.23 {+/-} 0.86); 63 of 69 low-accuracy translations (91.3%) occurred in the Nepali-to-English direction. Among 329 deviation tags, the most frequent were distortion of intended meaning (17.1%), overly formal or unnatural phrasing (14.7%), omission (14.2%), and addition of content (11.5%). Some fluent outputs substantially altered meaning or introduced information not expressed by the speaker. ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses. Accuracy was lower and more variable for Nepali-to-English translation; however, translation direction was confounded with input type because Nepali inputs were spontaneous and English inputs standardized, limiting conclusions about directional performance. These findings support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or participant understanding. As multimodal AI evolves, performance should be reevaluated across languages, real-world conditions, and model versions, with bilingual oversight and community partnership remaining central to responsible use.

9
Adapting Clinical Event Annotation to Dutch Primary Care: An Event Annotation Framework for Post-Acute Infection Syndromes

Mazzucato, S.; Leeuwenberg, A.; van Doorn, S.; van Rosmalen, J.; Slurink, I. A. L.

2026-08-22 health informatics 10.64898/2026.08.19.26360841 medRxiv
Top 0.2%
5.5%
Show abstract

Extracting clinical information from Dutch free-text medical notes requires language-specific annotation resources, yet Dutch primary care lacks a reusable event-annotation framework for infections, post-acute infection syndromes (PAIS), and related symptoms. We adapted the COVID-19 Annotated Clinical Text (CACT) framework to Dutch and applied it to GP notes for PAIS event extraction. The framework has three annotation layers: a DiagnosticExpression typology covering acute infections, post-acute syndromes, and relevant comorbidities; an eleven-subtype Evidence inventory grounded in Dutch primary-care testing practice; and explicit decision rules for the SOEP structure of Dutch general practitioner (GP) notes (Subjective, Objective, Evaluation, Plan), including the distinction between clinician hedging and patient-side hypotheticals. On a 200-note pilot, span-level F1 under the Lybarger criterion reached 0.51 [95% CI: 0.47, 0.55] across six core entities; restricted to spans both annotators noticed, conditional F1 reached 0.78 [0.75, 0.80], indicating that most disagreement stems from annotation coverage rather than label assignment. The adaptation illustrates how an English event-based clinical annotation framework can be extended to a new language and clinical setting, yielding a reusable resource for Dutch clinical NLP; which steps generalise beyond this case (CACT to Dutch primary care) and which are specific to Dutch or PAIS remain to be tested.

10
Drivers of Oncologist Preference of AI-Generated Literature Review in a Randomized Mixed-Methods Study

Bunning, B. J.; Weng, Y.; Wu, D. J.; Hui, G.; Hope, J. E.; Pandurangan, V.; Lopez, I.; Everett, S.; Chen, J. H.; Desai, M.

2026-08-27 health informatics 10.64898/2026.08.24.26361252 medRxiv
Top 0.2%
5.4%
Show abstract

Doctors increasingly rely on AI in the clinic, yet which report features make AI-generated responses useful and trustworthy remains unclear. In this randomized mixed-methods study, 34 oncology physicians provided 294 ratings of four blinded AI systems across five vignettes, alongside 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar references, an evidence-graded report adapted from OpenEvidence was rated significantly lower in overall utility than standard OpenEvidence (mean difference, -0.96; 95% CI, -1.26 to -0.66; P<.001). Qualitative analysis identified six themes and seven design requirements. Oncologists valued rapid orientation, evidence retrieval, and verification, preferring concise, scannable reports with quantitative outcomes, recognizable bolded guidelines, explicit uncertainty, and verifiable citations. Trust deteriorated with citation mismatch, buried provenance, evidence misclassification, overconfident recommendations, and poor organization. Evidence presented differently can alter perceptions of clinical utility and trust; accuracy alone is insufficient, and report design must also be empirically evaluated.

11
From Output Errors to Workflow Harm: A Practitioner-Audit Method for LLM-Mediated Research

Austria, D.; McCollister, B.; Lindsey, J. E.; Arowolo, M.; Okon, M.

2026-08-17 health informatics 10.64898/2026.08.13.26360414 medRxiv
Top 0.2%
5.4%
Show abstract

Objective. Formal large language model (LLM) evaluations score isolated prompts, but clinicians and health-informatics researchers meet model failures inside multi-step workflows where erroneous output can alter procedures or contaminate documents. We present TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI. Materials and Methods. A method paper with an empirical demonstration: 45 documentation-positive incidents recorded by one clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric, an error definition, a taxonomy crosswalk, and a Response-Audit Scorecard. Three reviewer-authors independently coded a 16-incident subsample; three vendor-blinded AI comparators applied the taxonomy to all 45 incidents. Results. Four categories tied as most frequent: verification failure, factual numerical error, tool-behavior misunderstanding, and citation or reference formatting (n=7 each). Four workflow-harm patterns recurred: procedural propagation, documentary contamination, trust-calibration disruption, and user-borne corrective burden, and one incident carried an estimated $2500 impact. Category agreement across three human reviewer-authors was low (Fleiss {kappa}=0.155), whereas three AI comparators agreed substantially (Fleiss {kappa}=0.632), suggesting taxonomy legibility under standardized conditions even where human judgment diverged. Discussion. Category assignment is comparatively legible, whereas severity and claimed-verification remain judgment-dependent. The claimed-verification gap is a measurable failure mode distinct from hallucination, sycophancy, and over-refusal. Conclusion. Practitioner audits with structured response scoring complement benchmarks by documenting workflow harm as an applied evaluation unit for clinical informatics and public-health work; this is a pilot that motivates, not estimates, error rates or cross-model comparisons.

12
A Pragmatic Randomized Trial of an EHR-Integrated Generative AI Chart Summarization Tool for Ambulatory Clinicians

Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.

2026-08-31 health informatics 10.64898/2026.08.26.26361496 medRxiv
Top 0.2%
5.4%
Show abstract

BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.

13
Are automated documentation-error judges fit to measure ambient AI scribes? A pre-registered, blinded human-validation study

Bergman, H. I.; Liu, V. N.; Austin, B.; Sanghera, R.

2026-08-17 health informatics 10.64898/2026.08.14.26360441 medRxiv
Top 0.2%
4.4%
Show abstract

Objectives Safety claims for ambient artificial intelligence (AI) scribes rest on automated judges that detect documentation errors and grade clinical risk. Expert reviewers are under-sensitive and disagree with one another, so no gold standard exists and validation cannot mean accuracy. We tested whether such judges are a defensible instrument: reproducible, within the envelope of expert disagreement, and non-differential across arms. Methods Pre-registered, blinded validation study nested in a multi-country simulation of ambient AI documentation (English setting), reported per GRRAS. Ten external clinicians independently adjudicated a stratified sample of 434 pipeline flags, retained and screen-discarded, blinded to note authorship, identification source, the pipeline's verdict and severity tier. Agreement used Gwet's AC1; proportions carry Wilson intervals. Three propositions were pre-specified: envelope parity, non-differential behaviour across arms, and concordance on consensus cases. Results All ten reviewers completed: 565 adjudications across 434 items, 131 of them double-rated. Inter-clinician agreement on genuineness was fair (raw 59%, 95% CI 50 to 67; AC1 0.24), leaving no human consensus to serve as truth. Judge-clinician agreement was 64% (95% CI 60 to 68), overlapping that interval. Behaviour was near-symmetric on contrast-critical metrics: kept-precision 74% for AI against 81% for clinician notes, and severity signed gap +0.06 against -0.09 tiers. One sub-metric was asymmetric: removed-confirmed 56% against 42%, so the screen over-removes more on clinician notes, a direction conservative to the parent contrast. On 77 consensus items the pipeline concurred on 70% (95% CI 59 to 79). Latent-class triangulation placed the genuine-error rate among flagged candidates at 68% (94% credible interval 48 to 83). Conclusions The judges behave as a consistent, near-non-differential, clinician-equivalent instrument. This licenses a directional AI-versus-clinician contrast under a non-differential misclassification argument, subject to its conditions. It is not a claim of accuracy, which moderate consensus concordance and fair reliability preclude, and the genuine-error rate is best reported as an interval.

14
Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients

Hamdan, M.; Harati, A.; Al-Bakheet, A.; Fuetterer, I.; Alshaer, I.

2026-08-06 surgery 10.64898/2026.08.04.26359718 medRxiv
Top 0.2%
4.2%
Show abstract

Objective: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease. Methods: We retrospectively analyzed 147 consecutive patients. Each case included clinical documentation and MRI presented as two composite PNG images. Two resident doctors and three multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the surgical level. Analyses used Cochran's Q, McNemar tests with Holm correction, and Bayesian methods. Results: LLMs achieved higher therapy-decision accuracy (66.0%-68.0%; 97-100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for residents versus 33.3%-41.1% for LLMs. Conclusion: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

15
Performance of an Ambient Generative AI Documentation Tool in a Linguistically Diverse Clinical Setting

Aldis, R.; Wang, S.; Sage, M.; Metzmaker, M.; Galvin, H.

2026-08-17 health systems and quality improvement 10.64898/2026.08.14.26360467 medRxiv
Top 0.2%
4.1%
Show abstract

Ambient artificial intelligence scribes are being increasingly used in healthcare to improve efficiency and reduce provider clinical documentation burden, yet their performance across linguistically diverse patient populations is not well characterized. We conducted a retrospective analysis of 54,160 outpatient encounters within a U.S. safety net health system to evaluate the performance of an artificial intelligence documentation tool in English and non-English clinical encounters, and in encounters where an interpreter or bilingual provider was present. Documentation performance was measured by the percentage of words in the final note that were generated by the ambient AI documentation tool and not edited by the provider. Associations between language factors and documentation performance were measured using Generalized Estimating Equations with exchangeable correlation structures to account for clustering of multiple encounters within unique patients. Univariable models were fitted to estimate the odds of adequate performance by language and interpreter modality, and a multivariable interaction model was used to evaluate within-language differences between bilingual providers and interpreter-mediated encounters. Non-English encounters were 21% to 25% less likely than English encounters to achieve the same performance threshold. There was no significant difference in generative documentation performance between interpreter-mediated and bilingual provider encounters. These findings underscore the importance of equity-focused evaluation and multilingual model refinement to ensure that artificial intelligence documentation benefits are distributed fairly across diverse patient populations.

16
Triage and Referral Behavior of Patient Facing Medical Artificial Intelligence Products

Margolis, S. J.; Maddipatla, N. V. S. K.; Ioannides, K. L. H.; Wisk, L. E.; Schriger, D. L.; Elmore, J. G.

2026-08-17 health informatics 10.64898/2026.08.15.26360518 medRxiv
Top 0.3%
3.9%
Show abstract

We evaluated nine patient-facing artificial intelligence products using 60 physician-developed standardized clinical cases and 540 multi-turn simulated patient encounters. Although overall triage accuracy showed no statistically significant difference across product categories, referral behavior differed substantially. Branded health AI products more frequently over-triaged low-acuity cases (28% vs 3% vs 2%) and recommended affiliated, fee-requiring clinical services. These findings suggest evaluation of patient-facing medical AI should assess referral behavior alongside overall triage accuracy.

17
A single-patient task exposes a failure of safety alignment in clinical language models

Gorenshtein, A.; Jia, E. L.; Omar, M.; Brook, O. R.; Ahmed, M.; Kruskel, J. B.; Barash, Y.; Klang, E.

2026-08-10 health informatics 10.64898/2026.08.07.26359822 medRxiv
Top 0.3%
3.4%
Show abstract

Safety alignment should persist while a language model performs a task. We tested whether a single-patient triage task suppressed a warning about a second patient. Each case centered on Patient 1; Patient 2's urgent problem appeared only in passing. Sixteen models saw each case twice: once as a general assistant and once while producing a triage record for Patient 1. As general assistants, models warned the caller in 87% of cases; under the task, they did so in 21%. Every model showed a significant decrease. Yet under the task, the record still mentioned Patient 2 in 76% of cases and recommended urgent care in 67%. Across 15 open-weight models, repeating the emergency-care instruction raised the warning rate only to 29%; moving the message-to-caller field to the top raised it to 36%. Current safety alignment did not reliably persist under task assignment.

18
Global Adoption of openEHR Clinical Data Repositories: A Vendor and Community Survey

Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.

2026-08-31 health informatics 10.64898/2026.08.27.26361529 medRxiv
Top 0.3%
3.4%
Show abstract

The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.

19
Perceived usability and usefulness of a clinical decision-support application among newly graduated physicians in rural areas: a mixed-methods study

De la Cruz-Torralva, K.; Diaz-Sanchez, P.; Escobar-Agreda, S.; Rojas-Mezarina, L.

2026-08-21 primary care research 10.64898/2026.08.18.26360759 medRxiv
Top 0.3%
3.4%
Show abstract

Mobile clinical-support applications can facilitate access to evidence-based information at the point of care, but evidence on their usability and perceived usefulness among newly graduated physicians working in health facilities with limited capacity is scarce. We assessed physicians experiences with BMJ Best Practice using a convergent mixed-methods study. All 81 eligible physicians assigned to rural facilities were invited; 32 enrolled and received application access and training. After three months, participants completed an online survey, and 23 reported using the application. Ten physicians reporting the highest consultation frequency were purposively selected for semi-structured interviews. Survey findings showed a predominantly favorable perception of usability: for most items, 70%-90% of participants agreed or strongly agreed with the statements assessed. Among users, 14 of 23 (60.9%) used the mobile application and 9 (39.1%) used the web version. Interviews indicated that participants valued rapid searches, organized and evidence-based information, and support for diagnostic reasoning, referral decisions, learning, and clinical confidence. Barriers included limited connectivity, difficulties searching in Spanish, automatic updates, challenges locating or using some calculators, and treatment information that was sometimes insufficiently specific. Most importantly, participants could not always implement recommendations because suggested medicines, diagnostic tests, or other resources were unavailable in their facilities. Mobile clinical-support applications may complement decision-making and learning among early-career physicians in rural primary care. However, their practical value depends not only on usability and evidence quality, but also on adaptation to users language, workflow, connectivity, and local service capacity.

20
Patient Perspectives on Potential Implementation of Coordinated Family Care Visits for Inherited Cardiovascular Disease

Draisin, E. R.; Badar, H.; Naik, H.; Platt, J.; Kaufman, B.; Salisbury, H.; Ison, H. E.

2026-08-07 cardiovascular medicine 10.64898/2026.08.05.26359830 medRxiv
Top 0.3%
3.3%
Show abstract

Introduction: Shared medical appointments (SMAs) are medical visits where multiple individuals are seen together in a group setting. For patients with inherited cardiovascular disease, where multiple family members often require ongoing cardiac care and screening, family SMAs may be particularly valuable as a tool to facilitate family communication and comprehension of their condition. This research aimed to identify patient perspectives on the potential benefits and challenges of family SMAs in comparison to an existing individual clinic model. Methods: Qualitative semi-structured interviews were conducted with adult family representatives. Each family had at least one family member seen at the adult and pediatric inherited cardiovascular disease clinics. Interview recordings were transcribed verbatim and inductively coded using a content analysis approach. Results: Sixteen families were interviewed in this study. The mean age of the family representative interviewed was 43.4 years ({+/-} 9.3 SD), and they were followed at Stanford Health Care for a mean of 7.3 years ({+/-} 4.2 SD). 81.2% (13/16) of families said they would find family SMAs beneficial. For interested families who consented to recorded interviews (n=12), benefits and challenges fell into two major categories: care quality and access and logistics. Interested families thought family SMAs would provide an added care quality benefit by increasing understanding among adults, children, and providers (83.3%, 10/12). Six of twelve participants interested in having family SMA visits felt there would be logistical/access-based benefits to this new model (50%, 6/12). Families also identified possible challenges with this model, such as less individualized care, potential privacy concerns, and concerns regarding the smoothness of the clinic process in coordinating a family SMA. Conclusion: The majority of families believed a family SMA model would provide added benefit to families with inherited cardiovascular disease, but requires thoughtful implementation and should be tailored to families? unique needs.